Papers with fine-grained evaluation framework

12 papers
TracSum: A New Benchmark for Aspect-Based Summarization with Sentence-Level Traceability in Medical Domain (2025.emnlp-main)

Copied to clipboard

Challenge: Existing evidence-based summarization tasks require tracing source evidence to assess their accuracy.
Approach: They propose a benchmark for traceable, aspect-based summarization that pairs summaries with sentence-level citations to enable users to trace back to the original context.
Outcome: The proposed benchmark can be used to evaluate document summarization with LLMs and human evaluations.
EXAMS: A Multi-subject High School Examinations Dataset for Cross-lingual and Multilingual Question Answering (2020.emnlp-main)

Copied to clipboard

Challenge: EXAMS is a benchmark dataset for cross-lingual and multilingual question answering for high school examinations.
Approach: They propose to use EXAMS to evaluate cross-lingual and multilingual question answering for high school examinations.
Outcome: The proposed model can be used to explore multilingual reasoning and knowledge transfer methods and pre-trained models in schools in different languages, which was not possible by now.
PSST: A Benchmark for Evaluation-driven Text Public-Speaking Style Transfer (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing methods to transfer text style focus on sentence-level data, limiting performance . current LLMs struggle to generate public speaking texts that align with human preferences .
Approach: They propose a task to transform official texts into public-speaking styles by analyzing real-world data.
Outcome: The proposed task aims to transform public speaking texts into public-speaking styles . the proposed framework analyzes characteristics and identifies problems of stylized texts .
FinGrAct: A Framework for FINe-GRrained Evaluation of ACTionability in Explainable Automatic Fact-Checking (2025.findings-emnlp)

Copied to clipboard

Challenge: despite the importance of actionability, no prior research has evaluated its effectiveness.
Approach: They propose a fine-grained evaluation framework that can access the web to assess actionability in AFC explanations.
Outcome: The proposed framework surpasses state-of-the-art evaluators in achieving highest correlation with human judgments while showing lowest egocentricbias.
A Unified Agentic Framework for Evaluating Conditional Image Generation (2025.acl-long)

Copied to clipboard

Challenge: Conditional image generation is a popular and personalization-oriented task, but there are challenges in developing task-agnostic, reliable, and explainable evaluation metrics.
Approach: They propose a unified agentic framework for comprehensive evaluation of conditional image generation tasks.
Outcome: The proposed framework achieves a high correlation with human assessments on seven prominent image generation tasks.
Reference Matters: Benchmarking Factual Error Correction for Dialogue Summarization with Fine-grained Evaluation Framework (2023.acl-long)

Copied to clipboard

Challenge: Current evaluations of FEC models that depend on factuality metrics are not reliable and detailed enough.
Approach: They propose a fine-grained evaluation framework that automatically evaluates FEC models on different error categories.
Outcome: The proposed evaluation framework compares models on different error categories and finds the best training modes and significant differences in the performance of existing models.
Dissecting Logical Reasoning in LLMs: A Fine-Grained Evaluation and Supervision Study (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing benchmarks that rely on final-answer accuracy fail to capture the quality of the reasoning process.
Approach: They propose a fine-grained evaluation framework that assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing.
Outcome: The proposed framework assesses logical reasoning across three dimensions: overall accuracy, stepwise soundness, and representation-level probing.
BAGELS: Benchmarking the Automated Generation and Extraction of Limitations from Scholarly Text (2025.findings-emnlp)

Copied to clipboard

Challenge: a growing number of scientific publications have limitations as a source of uncertainty.
Approach: They propose a computational architecture for extracting and generating limitations from scholarly papers using a novel Retrieval Augmented Generation technique.
Outcome: The proposed architecture extracts limitations from ACL, NeurIPS, and PeerJ papers and supplementes them with external reviews.
MT-RAIG: Novel Benchmark and Evaluation Framework for Retrieval-Augmented Insight Generation over Multiple Tables (2025.acl-long)

Copied to clipboard

Challenge: Existing studies on table-based reasoning focus on a single gold table, not multiple tables . a persistent demand for robust table understanding systems is resulting from the complexity of table data .
Approach: They propose a MT-RAIG Bench to evaluate systems on Retrieval-Augmented Insight Generation over Mulit-Tables.
Outcome: The proposed framework improves human quality judgments on the generated insights.
Memory-Driven Role-Playing: Evaluation and Enhancement of Persona Knowledge Utilization in LLMs (2026.findings-acl)

Copied to clipboard

Challenge: Existing models fail to recall and accurately apply designated persona knowledge without explicit cues . memory-driven role-playing paradigms are attracting significant interest .
Approach: They propose a memory-driven role-playing paradigm that frames persona knowledge as the LLM's internal memory store and a prompting architecture that guides structured memory retrieval and response generation.
Outcome: The proposed paradigm provides a comprehensive diagnostic for four-stage role-playing abilities across 12 LLMs.
Understanding GUI Agent Localization Biases through Logit Sharpness (2025.findings-emnlp)

Copied to clipboard

Challenge: Multimodal large language models often exhibit hallucinations that compromise reliability . despite promising performance, these models often display systematic localization errors .
Approach: They propose a framework that categorizes model predictions into four distinct types . they propose metric that evaluates alignment between semantic continuity and logits distribution .
Outcome: The proposed framework categorizes model predictions into four different types . it reveals nuanced failure modes beyond traditional accuracy metrics .
CMedCalc-Bench: A Fine-Grained Benchmark for Chinese Medical Calculations in LLM (2025.emnlp-main)

Copied to clipboard

Challenge: Existing medical NLP benchmarks focus on qualitative reasoning and textual comprehension, but lack of fine-grained evaluation of intermediate reasoning.
Approach: They propose a Chinese medical calculation benchmark that disentangles clinical entity extraction from numerical computation.
Outcome: The proposed framework disentangles clinical entity extraction from numerical computation, enabling systematic diagnosis of model deficiencies.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations